Conversation
hsliuustc0106
left a comment
There was a problem hiding this comment.
Reviewed commit 047fa89. Found one Windows window-selection issue below. Validation: 72 targeted tests passed; native desktop execution and real 4B inference were not run on this Linux host.
| screenshot=visual_model, | ||
| ), | ||
| session=_session(driver, flags.app), | ||
| session=_session(driver, flags.app_path or flags.app), |
There was a problem hiding this comment.
[P2] Pass the window title through Windows launch selection
--window-title reaches WindowEnv, but this session receives only the app name. On Windows, _session() calls CuaDriver.launch_app(), which rejects a launch result containing multiple visible windows before find_window() can apply the title filter. Consequently, an app with two open documents still fails even when exactly one matches the requested title. I reproduced this with a mocked driver response containing two windows: launch raises DriverError: launch_app: expected one on-screen window of 'Editor', got 2, and list_windows is never called. Pass the requested title through the launch path and filter the returned windows before enforcing uniqueness, while preserving the selected window's PID/window-ID binding.
There was a problem hiding this comment.
Thanks for catching this. I initially focused on macOS and missed this Windows launch case.
Fixed in cf49ee6. --window-title now filters the launch results before the single-window check. The selected window stays bound to its PID and window ID.
Added regression tests with mocked Windows driver responses for title selection, no match, duplicate matches, and window binding. Full suite: 509 passed, 41 skipped. Ruff, ty, and CLI/MCP smoke checks passed.
What this enables
The desktop agent can use a local Cua-S1 4B model to choose a click target from the current window's screenshot. This supports controls that have no accessibility element.
The task defines named candidate points with
--pixel-target. The model receives the screenshot and selects one of those points. Each click carries the window target and capture ID, so the driver can reject stale captures and invalid coordinates.The optional
cua-four-bextra supports the text and multimodal adapters through the existing decision-model interface.CUA_S1_VARIANT=4bselects it; Nano remains the default for--model cua. Screenshot bytes are passed separately from the serialized observation state.Implements RFC #30. Related features: #14 and #15.
macOS results
Tested on macOS 26.2 arm64 with Cua Driver 0.30.1. The native app draws Save and Cancel without accessibility children and swaps their positions on reset.
The model selected right, left, right, left as Save moved. All four selections were checked against the app's output file. Median model decision time was 7.369 s. Episode times exclude model loading; the model run used no fixed plan or hosted model API calls.
An intentionally wrong-point check was rejected and recorded Cancel in the output file.
Tests
Added
tests/test_decision_models_cua_four_b.pyandtests/test_desktop_visual.py, and extended the model factory tests. Coverage includes option mapping, invalid model outputs, image delivery, temporary-image cleanup, capture-bound clicks, dry runs, window selection, and separate driver sessions.Windows regression tests also cover title-based launch selection, missing or duplicate matches, and PID/window-ID binding using mocked driver responses.
Core + dev suite: 509 passed, 41 skipped. Ruff, ty, lock check, package build, shell syntax, and CLI/MCP smoke passed.
How to run
Requires macOS,
uv, Swift command-line tools, and Cua Driver with Accessibility and Screen Recording permissions. Keep the Mac unlocked and close any older S1A visual fixture. Run from the current project directory:The base model and adapter download on first use.
CUA_S1_BASE_MODELandCUA_S1_CHECKPOINTcan point to downloaded model directories.